gpt-oss-120b · Server
gpt-oss-120b / Server / Tokens/s
v6.0|closed|gpt-oss-120b|Server|Tokens/s|official-system
Metric: Performance_Result (Tokens/s) · winning: higher · 28 ranking-eligible rows · vendors 3 · families 7 · never mixed with other slices
All systems (official)
| Rank | Performance | Unit | System | Submitter | Accelerator | Source |
|---|---|---|---|---|---|---|
| 1 | 1,096,770.0000 | Tokens/s | Nebius GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) | Nebius | NVIDIA GB300 | source |
| 2 | 1,072,250.0000 | Tokens/s | NVIDIA GB300 NVL72 (72x GB300-288GB_aarch64, TensorRT) | NVIDIA | NVIDIA GB300 | source |
| 3 | 956,523.0000 | Tokens/s | CoreWeave GB300 NVL72 (64x GB300-288GB_aarch64, TensorRT) | CoreWeave | NVIDIA GB300 | source |
| 4 | 900,054.0000 | Tokens/s | AMD MI355X Custom Cluster | AMD | AMD Instinct MI355X 288GB HBM3e (x94) | source |
| 5 | 899,218.0000 | Tokens/s | NVIDIA GB200 NVL72 (72x GB200-186GB_aarch64, TensorRT) | NVIDIA | NVIDIA GB200 | source |
| 6 | 743,550.0000 | Tokens/s | CoreWeave GB200 NVL72 (64x GB200-186GB_aarch64, TensorRT) | CoreWeave | NVIDIA GB200 | source |
| 7 | 110,655.0000 | Tokens/s | Cisco UCS C880A M8 (8x NVIDIA B300-SXM-270GB, TensorRT) | Cisco | NVIDIA B300-SXM-270GB | source |
| 8 | 106,397.0000 | Tokens/s | G894-SD3-AAX7 | GigaComputing | NVIDIA B300-SXM-270GB | source |
| 9 | 100,687.0000 | Tokens/s | ThinkSystem SR680a V4 (8x B300-SXM-270GB, TensorRT) | Lenovo | NVIDIA B300-SXM-270GB | source |
| 10 | 100,656.0000 | Tokens/s | XA NB3I-E12 | ASUSTeK | NVIDIA B300-SXM-270GB | source |
| 11 | 100,437.0000 | Tokens/s | Nebius B300 n1 (8x B300-SXM-270GB, TensorRT) | Nebius | NVIDIA B300-SXM-270GB | source |
| 12 | 100,328.0000 | Tokens/s | NVIDIA DGX B300 (8x B300-SXM-270GB, TensorRT) | NVIDIA | NVIDIA B300-SXM-270GB | source |
| 13 | 87,444.2000 | Tokens/s | Nebius B200 n1 (8x B200-SXM-180GB, TensorRT) | Nebius | NVIDIA B200-SXM-180GB | source |
| 14 | 82,136.1000 | Tokens/s | smci355-ccs-aus-m09-17 (AS -4126GS-NMR-LCC) | AMD | AMD Instinct MI355X 288GB HBM3e | source |
| 15 | 80,089.6000 | Tokens/s | G893-ZX1-AAX4 | GigaComputing | AMD Instinct MI355X 288GB HBM3e | source |
| 16 | 78,603.5000 | Tokens/s | PowerEdge XE9785L (8x MI355X) | Dell | AMD Instinct MI355X 288GB HBM3e | source |
| 17 | 78,186.3000 | Tokens/s | BM.GPU.MI355X.8 | ORACLE | AMD Instinct MI355X 288GB HBM3e | source |
| 18 | 75,354.9000 | Tokens/s | HPE ProLiant Compute XD685 (8x AMD Instinct MI355X 288GB, vLLM) | HPE | AMD Instinct MI355X 288GB HBM3e | source |
| 19 | 71,588.1000 | Tokens/s | LLM-D v0.5.0,Openshift 4.20.12,NVIDIA 8xB200-SXM-180GB | RedHat | NVIDIA B200-SXM-180GB | source |
| 20 | 53,462.7000 | Tokens/s | NVIDIA GB300 NVL72 (4x GB300-288GB_aarch64, TensorRT) | Lambda_SIT | NVIDIA GB300 | source |
| 21 | 24,103.2000 | Tokens/s | LLM-D v0.5.0,OpenShift 4.21,IBM cloud gx3d-160x1792x8h200,NVIDIA 8xH200-SXM-141GB | RedHat | NVIDIA H200-SXM-141GB | source |
| 22 | 17,737.6000 | Tokens/s | HPE ProLiant Compute DL380a Gen12 (10x NVIDIA RTX PRO 6000 Blackwell Server Edition, TensorRT) | HPE | NVIDIA RTX PRO 6000 Blackwell Server Edition | source |
| 23 | 14,973.1000 | Tokens/s | Nebius GB300 NVL72 (4x GB300-288GB_aarch64, TensorRT) | Nebius | NVIDIA GB300 | source |
| 24 | 14,258.9000 | Tokens/s | HPE ProLiant Compute DL380a Gen12 (8x NVIDIA RTX PRO 6000 Blackwell Server Edition, TensorRT) | HPE | NVIDIA RTX PRO 6000 Blackwell Server Edition | source |
| 25 | 6,687.2500 | Tokens/s | HPE ProLiant DL385 Gen11 (4x NVIDIA RTX PRO 6000 Blackwell Server Edition-450W, TensorRT) | HPE | NVIDIA RTX PRO 6000 Blackwell Server Edition | source |
| 26 | 951.6740 | Tokens/s | 1-node-4x-BMG-B70 | Intel | Intel(R) Arc Pro(R) B70 | source |
| 27 | 884.2380 | Tokens/s | 1-node-4x-BMG-Pro-B60-Dual | Intel | MS-Intel Arc Pro B60 Dual 48G Turbo | source |
| 28 | 452.1910 | Tokens/s | 1-node-4x-BMG-B60 | Intel | Intel(R) Arc Pro(R) B60 | source |